The Telephone Game: Vanishing Gradients
We just learned that to train an RNN, we use BPTT (Backpropagation Through Time). We take the final error and pass it backward through every step of the unrolled network to adjust the weights.
But what if the sentence is 100 words long?
Passing math backwards through 100 steps creates a massive, catastrophic mathematical glitch. This glitch is the reason Vanilla RNNs are almost never used in the real world anymore.
Let's look at the two ways this glitch ruins our AI: Vanishing and Exploding gradients.
1. The Vanishing Gradient (The Fading Whisper)
Have you ever played the game of "Telephone"? You whisper a secret to your friend, they whisper it to the next person, and so on down a line of 20 people. By the time it reaches the last person, the secret is completely lost or faded away.
This is exactly what happens to our error signal in an RNN!
Every time the error takes a step back in time, it gets multiplied by the network's weights. Because of how we initialize neural networks, these weights are usually small decimals (like 0.5).
If you multiply a number by 0.5 over and over again, look what happens:
- Step 100:
Error = 1.0 - Step 99:
1.0 * 0.5 = 0.5 - Step 98:
0.5 * 0.5 = 0.25 - Step 97:
0.25 * 0.5 = 0.125 - ...
- Step 1:
0.0000000000000000001
The Result: By the time the error reaches the beginning of the sentence, it has vanished into nothingness. The AI completely fails to learn how the beginning of the sentence affects the end of the sentence. (Goldfish memory!)
2. The Exploding Gradient (The Screaming Megaphone)
Now imagine the opposite. Instead of small decimals, what if our network weights happen to be slightly larger than 1 (like 1.5)?
Instead of fading away, the math acts like a microphone placed too close to a speaker. The sound loops back, gets amplified, and creates a deafening screech!
- Step 100:
Error = 1.0 - Step 99:
1.0 * 1.5 = 1.5 - Step 98:
1.5 * 1.5 = 2.25 - ...
- Step 1:
999,999,999,999,999.0
The Result: The error becomes so massively huge that it completely breaks the computer's memory. The AI's brain turns to mush, and the training immediately crashes.
The Solution?
Vanilla RNNs are broken. If the math is less than 1, it vanishes. If it's more than 1, it explodes. It is almost impossible to balance perfectly.
We need a completely new architecture. We need an AI that doesn't just blindly multiply numbers at every step. We need an AI that has a Brain—one that can choose what to remember, and what to forget.
Next Up: Welcome to Chapter 4, where we introduce the savior of Recurrent Neural Networks: the LSTM (Long Short-Term Memory)!